Principal value analysis (PCA)
Lecture 33
Applications of SVD
$$ % Colors
% Coordinate vectors and matrices
% Common sets
% Abstract vector symbols
% Norms / absolute value
% Optional: dot product spacing (looks nicer in slides)
% Operators $$
Dimension reduction
- One important interpretation of SVD is that larger singular values encode the main structure (or important core features) of the data, while smaller singular values often encode fine details or noise
- This is analogous to Fourier approximation: low-frequency components capture the overall shape (coarse features), while high-frequency components capture fine details or noise
- In particular, by dropping small singular values (and their corresponding singular vectors), we can reduce the dimension of the data while preserving its essential structure
- This idea leads to low-rank approximation: instead of using the full SVD \[ A = U\Sigma V^T, \] we approximate it by keeping only the top \(k\) singular values: \[ A \approx U_k \Sigma_k V_k^T \] where \(k\) is much smaller than the rank of \(A\)
- This is called the truncated SVD and provides the best rank-\(k\) approximation of \(A\) (in least squares sense)
- Practical interpretation:
- Data compression (store fewer numbers)
- Noise reduction (discard small singular values)
- Feature extraction (keep dominant patterns)
Image source: toward data science
Digit recognition
- Demonstration for digit recognition through SVD
- A handwritten digit image can be viewed as a point in a high-dimensional vector space.
- For MNIST-style data, each image has \(28\times 28 = 784\) pixels, so one image becomes a vector in \(\mathbb{R}^{784}\).
- If we collect many labeled images of digits \(0,1,\dots,9\), those images are not scattered randomly in \(\mathbb{R}^{784}\). Digits with the same label tend to share common geometric patterns such as loops, vertical strokes, and slanted lines.
- The goal is to discover those dominant patterns automatically and use them to compress and classify the images.
Detailed explanation
- Suppose we have \(K\) labeled training images \(x_1,\dots,x_K\in\mathbb{R}^{784}\), where each label is one of the digits \(0,1,\dots,9\).
- Put these images into one matrix: \[ A=\begin{bmatrix}x_1 & x_2 & \cdots & x_K\end{bmatrix}\in\mathbb{R}^{784\times K}. \]
- Each column of \(A\) is one image written as a long vector.
- Even though the ambient space has dimension \(784\), the data often lie near a much lower-dimensional set, because handwritten digits share strong structural similarities.
- This is exactly why SVD is useful: it finds the main directions in which the dataset varies.
SVD-based digit recognition
To actually recognize digits, we build one SVD model for each digit \(d=0,1,\dots,9\).
For each digit \(d\), collect all training images with that label: \[ A_d=\begin{bmatrix} x^{(d)}_1 & x^{(d)}_2 & \cdots & x^{(d)}_{K_d} \end{bmatrix}\in\mathbb{R}^{784\times K_d}. \]
Step 1: Compute the mean image for digit \(d\) \[ \mu_d=\frac1{K_d}\sum_{j=1}^{K_d}x^{(d)}_j. \]
Step 2: Center the data \[ \widetilde A_d= \begin{bmatrix} x^{(d)}_1-\mu_d & x^{(d)}_2-\mu_d & \cdots & x^{(d)}_{K_d}-\mu_d \end{bmatrix}. \]
Step 3: Compute SVD \[ \widetilde A_d = U_d\Sigma_dV_d^T. \]
Step 4: Keep top \(r\) singular vectors \[ U_{d,r}=\begin{bmatrix}u^{(d)}_1 & \cdots & u^{(d)}_r\end{bmatrix}. \]
This defines a low-dimensional subspace for digit \(d\).
Now given a new image \(y\in\mathbb{R}^{784}\):
For each digit \(d=0,\dots,9\):
Step 5: Center using digit mean \[ z_d = y-\mu_d. \]
Step 6: Project onto subspace \[ p_d = U_{d,r}U_{d,r}^T z_d. \]
Step 7: Reconstruct \[ \widehat y_d = \mu_d + p_d. \]
Step 8: Compute reconstruction error \[ e_d = \|y-\widehat y_d\|_2 = \|z_d - U_{d,r}U_{d,r}^T z_d\|_2. \]
Step 9: Compare all errors \[ e_0, e_1, \dots, e_9 \]
Step 10: Predict label \[ \mathrm{label}(y) = \arg\min_{d} e_d. \]
Deep Learning
- This classical SVD/PCA approach is still very valuable because it is simple, interpretable, and mathematically transparent.
- It shows us how image data can be compressed, denoised, and organized geometrically.
- However, it is still a linear model: it assumes the important structure of the data can be captured by a linear subspace.
- Real handwriting can vary in more complicated ways — rotation, thickness, translation, curvature, and writing style — so purely linear models have clear limitations.
- In modern computer visions, digit-recognition examples are commonly built with neural networks.
- For example: digit recognition demo
Principal component analysis (PCA)
Motivation
- Recall that if \[ A=[u_1\ \cdots\ u_m], \qquad B=[w_1\ \cdots\ w_n], \] then \[ (A^TB)_{ij}=u_i\cdot w_j. \]
- Therefore, \(A^TB\) stores all pairwise dot products between the columns of \(A\) and the columns of \(B\).
- In particular, \(A^TA\) it measures how similar the samples are to one another.
- If the centered data vectors are stored as columns of \(A\), then \(A^TA\) describes sample-to-sample similarities.
- PCA asks a geometric question: in which directions does the data vary the most? The principal components are the orthogonal directions that capture that variation most efficiently.
- Once we find those directions, we can rotate the coordinate system to align with the data and often keep only a few coordinates without losing much information.
Example
- This article gives an example of PCA.
- Imagine each data point is a pair \((x,y)\), where \(x\) represents happiness and \(y\) represents achievement.
- If the point cloud is stretched along a diagonal direction, then most of the variation happens along that diagonal.
- The first principal component is the unit vector pointing in that dominant direction.
- The second principal component is perpendicular to the first and captures the largest remaining variation.
- So PCA can be interpreted as a rotation of coordinates: instead of using the original axes, we use new axes that match the actual geometry of the data cloud.
- This is why PCA is useful for visualization and dimensionality reduction.
Principal component analysis
- Let \(X\in\mathbb{R}^{n\times K}\) be the centered data matrix, with samples stored as columns.
- The first principal component is the unit vector \(u_1\) that maximizes \[ \sum_{j=1}^K (u^Tx_j)^2=\|u^TX\|^2. \]
- Since \[ \|u^TX\|^2=u^T(XX^T)u, \] the first principal component is an eigenvector of \(XX^T\) corresponding to the largest eigenvalue.
- The second principal component is defined the same way, with the extra condition that it must be orthogonal to the first.
- Continuing in this way produces an orthonormal sequence of directions ordered by decreasing importance.
- If \[ X=U\Sigma V^T, \] then the principal directions are exactly the columns of \(U\), and the variance explained by the \(i\)-th component is proportional to \(\sigma_i^2\).
- Projecting onto the first \(r\) principal components gives the best rank-\(r\) linear approximation to the centered data in the least-squares sense.
Visualization
- This is another nice example of visualizing PCA.
- It shows PCA as a rotation toward the directions of maximum variance.
- This is often the fastest way to build geometric intuition before doing the algebra.